Back

Journal of Computational Chemistry

Wiley

Preprints posted in the last 90 days, ranked by how well they match Journal of Computational Chemistry's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Glycine molecule radical: Predicted properties and dipeptide formation

Synak, J.; Blazewicz, J.

2026-07-10 bioinformatics 10.64898/2026.07.07.736934 medRxiv
Top 0.1%
6.8%
Show abstract

Numerous advances in quantum and computational chemistry over the last decades, well as the development of computer science, allowed utilisation of more precise and complex models, which can be now applied to much bigger systems than in the past. The authors used Gaussian, coupled with theoretical methods, to predict a new way of peptide bond formation, which could have taken place in prebiotic conditions. To better tackle this difficult task, the properties of substrates (glycine-derived radicals) were extensively analysed, using the aforementioned tool - Gaussian, paired with taking resonance and hybridisation into account, to better understand the stereochemistry and the very nature of processes taking place. The result is a series of reactions, which without any sophisticated catalysts and with relatively low energy thresholds ({inverted exclamation}20 kcal/mol) can lead to formation of dipeptides (and further, oligopeptides). The authors also hope, the other predicted properties of the investigated molecules can be of use to any researcher, who would like to utilise them in their experiments. Author summaryOur goal was to investigate a way first peptide bonds in prebiotic conditions could have been formed. This is an extremely important step in research into the beginning of life on Earth. We found a very promising series of reactions, which uses atomic hydrogen as its only catalyst and confirmed our expectations with theoretical calculations, using Gaussian. There are two radicals derived from glycine, which perform major roles in the process, so we investigated their properties with Gaussian and verified that the results are in agreement with our own theoretical considerations. This involved checking for possible geometric isomers and conformers and creating models which could explain their properties. We are well aware that such calculations have limitations and there is no model, which is 100% accurate, so our results should be further confirmed by empirical data in the future. However, we still to be as thorough as possible in how we approached the subject.

2
Integrating the MARTINI2 Coarse-Grained Force Field into HADDOCK3 for Faster Modelling of Large Biomolecular Complexes

Versini, R.; Reys, V. G. P.; Kravchenko, A.; Honorato, R. V.; Bonvin, A. M. J. J.

2026-04-27 bioinformatics 10.64898/2026.04.25.720800 medRxiv
Top 0.1%
5.2%
Show abstract

The integration of coarse-grained (CG) approaches into docking workflows offers a powerful strategy for modelling large biomolecular assemblies with reduced computational costs. We present here the implementation of the MARTINI2 coarse-grained force field into the HADDOCK3 integrative modelling platform. This development enables the use of the CG representations and parameters within HADDOCK3 for efficient sampling and scoring of large protein-protein complexes. The implementation takes advantage of the modular and flexible architecture of HADDOCK3, allowing a seamless combination of MARTINI2 representation with the various modules. Conversion from and to all-atom models is integrated into the coarse-grained modelling workflow. The performance of the protocol is first assessed on protein-protein and protein-DNA benchmarks and then illustrated on a few representative large-scale systems, demonstrating a significant reduction in computational costs while maintaining biologically relevant accuracy.

3
The Quantum Environment in Cryptochrome Enhances Light Absorption of FAD

Wieners, L.; Garcia, M. E.

2026-04-28 biophysics 10.64898/2026.04.24.720615 medRxiv
Top 0.1%
4.8%
Show abstract

The light absorption of the protein cryptochrome and its chromophore FAD is important for the regulation of circadian rhythms and in some species for sensing magnetic fields. To compute the absorption spectrum of chromophore, typically only a small region is treated quantum-mechanically due the high computational cost of spectroscopic calculations. We present a formalism that allows a quantum-mechanical treatment of not only the chromophore but also the neighbouring amino acids which differ from species to species. This is achieved by using the real-time time-dependent Hartree-Fock method. This method allows extending the quantum domain from typically only a few dozen atoms up to around 1,200 atoms for the largest calculations. The presented framework allows the treatment of neighbouring tryptophan residues or the cofactor molecule MTHF in the same calculation and allows to extract information of which regions absorb light depending on wavelength. The presented results also show that the environment around the chromophore FAD amplifies the light absorption in cryptochrome.

4
SuBMIT: A Software Toolkit for Facilitating Simulations of Coarse-Grained Structure-Based Models of Biomolecules.

Prakash, D. L.; Banerjee, A.; Gosavi, S.

2026-05-20 biophysics 10.64898/2026.05.18.725912 medRxiv
Top 0.1%
4.4%
Show abstract

Coarse-grained structure-based models (CG-SBMs; or G[o] models) are simplified potential energy functions of biomolecules or biomolecular complexes that encode their structure. Molecular dynamics simulations of such SBMs have been successfully used to study long time-scale dynamics such as protein and RNA folding, and large conformational transitions of biomolecular complexes. SBMs have several advantages: (1) Their MD simulations are computationally inexpensive, making extensive sampling easily accessible to many researchers. (2) They are easy to modify and can be adapted for the specific biomolecular problem that needs to be investigated. However, the force-fields of SBMs are not usually included in commonly used biomolecular simulation packages resulting in a barrier to their use. Here, we present SuBMIT (Structure Based Models Input Toolkit; https://github.com/sglabncbs/submit), a toolkit for generating coarse-grained SBM input files for performing MD simulations with GROMACS and OpenMM/OpenSMOG. Simulations whose input files can be generated using the different flavors of CG-SBMs present in SuBMIT include the folding and conformational ensembles of proteins with intrinsically disordered regions, 3D-domain-swapping in proteins and the dynamics of RNA-protein assemblies (e.g., simple RNA viruses).

5
A Pipeline for Solving Edge-Matching Puzzles and Their Implications for Protein Folding

Seifer, S.

2026-05-24 biophysics 10.64898/2026.05.23.727379 medRxiv
Top 0.1%
4.0%
Show abstract

Progress in quantum computation offers new opportunities for addressing longstanding combinatorial challenges. One such challenge is the Eternity II edge-matching puzzle, consisting of 256 tiles, which has resisted solution despite extensive community effort. The computational complexity of this NP-complete problem exceeds the capacity of current quantum annealing processors but lies within reach of hybrid quantum-classical solvers. Testing a quadratic unconstrained binary optimization (QUBO) model of a puzzle on a D-Wave hybrid solver demonstrates a complete solution only for puzzle instances up to 64 tiles. Simulated quantum annealing fails on this benchmark, whereas an original classical heuristic, "nucleation with deduction", succeeds. To approach the full Eternity II puzzle, I developed a MATLAB package that integrates multiple quantum and classical approaches, including neural-network transformers and gradient-based refinement. A multistage computation pipeline is demonstrated successfully on a puzzle comparable in complexity to Eternity II and with a known solution, based on multiple hybrid optimization steps with both "hard" and "soft" constraint formulations, identification of persistent substructures, and a final classical refinement stage. The resulting optimization problem involves [~]100,000 logical variables and requires partial initialization. Intriguingly, solving this puzzle mirrors the "end game" of protein folding, a process that nature completes in mere fractions of a second, seemingly defying expectations set by the Levinthal paradox. The prospect of predicting protein structure by quantum annealing is reviewed in light of these results.

6
Solvent-buffer effects in molecular dynamics simulations of nucleic acids

Baghel, N.; Shrivastava, P.; Mehra, R.

2026-07-06 biophysics 10.64898/2026.07.05.736650 medRxiv
Top 0.1%
3.6%
Show abstract

Molecular dynamics simulations of nucleic acids are performed using a solvent-buffer distance of 10 [A] between the solute surface and the simulation box boundary. Although this cell size has been extensively explored in protein simulations, its implications for nucleic acid dynamics are not well understood. Nucleic acids are elongated, highly charged, and flexible structures with hydration and dynamical properties distinct from those of proteins and therefore, they may require different solvent-layer considerations in simulations. In this study, we investigated the effect of simulation cell size on nucleic acid dynamics by simulating a 30-base-pair double-helical nucleic acid structure and its two single-stranded forms using solvent-buffer distances of 3, 5, 10, 15, and 20 [A]. Smaller cells may impose restricted hydration, molecular crowding, and periodic image interactions. However, larger cells provide solvent space for conformational relaxation. A total of 45 s of molecular dynamics simulations were performed (3 structures x 5 cell sizes x 3 replicates x 1 s). Our results show that while the commonly used 10 [A] buffer may be sufficient to maintain the stability of the double-stranded nucleic acid, larger cells are required to capture the conformational dynamics of single-stranded structures. In both, increasing the cell size to 15 or 20 [A] enables broader conformational sampling. The first hydration shell exhibits reduced crowding in the 20 [A] cell, consistent with more relaxed conformations. At larger cell sizes, single-stranded nucleic acids adopt compact, self-associated conformations for stability. Together, this study presents physical insight into how simulation cell size and solvent environment influence nucleic acid dynamics.

7
AlphaUnfold: Probing Potential Unfolding and Structural Fragility in AlphaFold3 Models via Short-Time High-Pressure MD

Pegado, F. J. d. O.; Ortega, J. M.; Silva, J. R. P.

2026-04-26 bioinformatics 10.64898/2026.04.22.720259 medRxiv
Top 0.1%
3.3%
Show abstract

We developed AlphaUnfold, an automated pipeline that couples AF3 predictions with short-time (5 ns) high-pressure Molecular Dynamics (MD) using NAMD3. By subjecting models to baric stress, AlphaUnfold acts as a dynamic "stress-test" to identify structural fragility and potential unfolding. Testing a diverse set of proteins revealed a significant inverse correlation between average pLDDT and Root Mean Square Deviation (RMSD) after MD, indicating that lower confidence translates to rapid structural drift. Furthermore, domains with low local pLDDT consistently exhibited high Root Mean Square Fluctuation (RMSF), a behavior also observed in 200 ns simulations under standard pressure, pinpointing specific metastable areas. AlphaUnfold thus provides a viable, computationally efficient framework for assessing the biophysical robustness of AI-generated models, offering an "experimental-like" validation that ensures more reliable downstream applications in structural biology. MotivationAlphaFold3 (AF3) provides high-accuracy protein models characterized by the Predicted Local Distance Difference Test (pLDDT). However, these static predictions may harbor "not well-forged" regions lacking thermodynamic resilience. There is a critical need for rapid computational protocols to validate structural integrity beyond static confidence scores. AvailabilityGitHub: https://github.com/pegados/pipeline_AlphaUnfold Supplementary informationSupplementary data are available at http://biodados.icb.ufmg.br/alphaunfold Contacte-mail fabio, silva-jrp.miguel@ufmg.br

8
Improving All-Atom Molecular Dynamics Models for Quantitative Prediction of Nanopore Blockade Current

Liu, J.; Rodriguez, C.; Chen, M.; Aksimentiev, A.

2026-06-14 biophysics 10.64898/2026.06.12.731905 medRxiv
Top 0.1%
2.5%
Show abstract

All-atom molecular dynamics has become an indispensable tool in development of nanopore sensors of biological information. In a typical nanopore experiment, measurements of ionic current flowing through a nanopore report on the chemical structure of biomolecules that pass through the nanopore. Such experiments alone are often insufficient to relate the structure of the biomolecules to the ionic current modulations. The molecular dynamics method can establish such a relationship directly through a brute force simulation under applied electric field. Here, we examine the ability of molecular dynamics force fields to reproduce experimentally measured nanopore blockade currents produced by single-stranded DNA. Our simulations show that none of the "off the shelf" force fields (CHARMM36, AMBER Parmbsc1 and DES-AMBER) is capable of reproducing experimental data with the desired level of accuracy. To improve the accuracy, we examined and refined interactions between ions, protein nanopores and DNA, guided by experiments designed specifically for this purpose. Ultimately, the introduction of surgical corrections to non-bonded interactions within the CHARMM36 force field produced a favorable agreement between simulation and experiment. This refined parameterization, initially developed for nanopore sensing simulations, may have broader applications in computational studies of DNA-protein systems.

9
Comparative Analysis of Relative Ligand Binding Free Energy Simulation Methods: Amber-TI, GROMACS-NETI, OpenMM-FEP, and BLaDE-MSLD

Lee, H.; Kim, I.; Kim, S.; Bae, M.; Jeong, B.; Kim, S.; Jo, S.; Lee, J.; Im, W.

2026-04-24 biophysics 10.64898/2026.04.22.720125 medRxiv
Top 0.1%
2.4%
Show abstract

Structure-based drug design has become increasingly important in the pharmaceutical industry for accelerating the discovery of effective drug candidates. In particular, ligand binding free energy serves as a critical metric for predicting drug efficacy during the key stages of hit discovery and lead optimization. Continuous progresses have been made in the prediction of ligand binding free energies, but direct comparisons of different methods using the same force field remain challenging due to their unique implementations into different simulation engines. In this study, we present a direct comparison of four popular methodologies (Amber-TI, GROMACS-NETI, OpenMM-FEP, and BLaDE-MSLD) for calculating relative binding free energies ({Delta}{Delta}Gbind) with the same Amber protein and ligand force fields using MolCube Alchemical Free Energy Simulator (MolCube-AFES), which provides an input generation workflow to support {Delta}{Delta}Gbind calculations of all four methods. We used 80 alchemical transformations (among the JACS benchmark set by Wang et al.) and two additional applications to compare the predicted {Delta}{Delta}Gbind from the four methods against experimental measurements. All four methods reproduced experimentally observed trends with most transformations within {+/-}2 kcal/mol from experiments and show broadly comparable accuracy with no statistically significant performance differences across the benchmark dataset. These results demonstrate that MolCube-AFES enables controlled, cross platform benchmarking and show that all four different alchemical free energy methods deliver statistically equivalent accuracy, with method selection guided by workflow requirements such as throughput, portability, and perturbation network design rather than expected differences in performances.

10
AI-derived Protein Structures Validation: AlphaFold2 Models in the Twilight Zone

Griffin, P.; Deganutti, G.; Jadeja, K.; Idigbe, C.; Pipito', L.; Mejuto, L.; Ng, C. P.; Peck, S.; Greaves, J.; Reynolds, C. A.

2026-05-12 bioinformatics 10.64898/2026.05.12.724499 medRxiv
Top 0.1%
2.4%
Show abstract

In any field, unquestioningly accepting artificial intelligence (AI) results should be considered bad practise. Here, we devised a comparative modelling-based strategy for validating protein structures that exploits the well-known observation that protein folds are far more conserved than protein sequences. We identify proteins with a similar fold to the AlphaFold-generated query protein and determine their structural alignment to the query. The hypothesis is that if the sequence alignment coincides with the structural alignment, then the structure is validated. The strategy is implemented on a helix-by-helix and strand-by-strand basis using a multi-template pairwise local profile alignment method that works well into the twilight zone. The method is illustrated by application to the transmembrane transporter PEPT1, for which the structure is known, and the S-deacylases ABHD13 and ABHD16A, for which only AI-generated models exist. ABHD16A is particularly challenging because a sequence alignment search with BLASTp does not reveal any structural homologues and therefore requires work with extremely remote homologues; however, both models are validated through this strategy and are stable during classical molecular dynamics simulations. The ability of the strategy to identify errors is assessed with reference to misaligned ABHD13 models and misfolded decoy proteins.

11
Protein hydration and druggability

Panasenko, S.; Khorev, V.; Petukhov, M.

2026-07-08 biophysics 10.64898/2026.07.06.736750 medRxiv
Top 0.1%
2.1%
Show abstract

A priori assessment of target proteins' druggability remains an unsolved problem in the field of drug development. The empirical approaches widely used to solve this problem demonstrate low efficiency. In this work, we investigated the factor of hydration of a representative set of 65 evolutionarily and structurally unrelated human enzymes in a water environment. This factor depends only on the structure of the proteins, and not on the physical and chemical properties of any potential ligands. The results show that, unlike the widely used approaches based on calculations of the accessible surface area (ASA), the content of low-entropy water molecules (LEW) in the active sites of human enzymes is systematically higher than that in other areas of their surface, including inactive cavities. Optimal criteria and a step-by-step procedure for identifying protein ligand binding sites are proposed. The proposed approach, based on the calculation of the LEW content in the first hydration layer of potentially interesting target proteins, makes it possible to evaluate their medicinal suitability even before the development of any ligands. The article also presents the results of a comparative analysis of experimental Raman spectroscopy data and the results of molecular dynamics simulations of water hydrogen bonds using three widely used water models (TIP3P, OPC3, and TIP5P) and standard algorithms for calculating hydrogen bond networks.

12
Collinearity of Decomposed Energy Terms in MM-GBSA Binding Free Energy Calculations

Sevim, A.; Kocak, A.

2026-06-29 biophysics 10.64898/2026.06.24.734195 medRxiv
Top 0.1%
1.9%
Show abstract

The molecular mechanics-generalized Born surface area method (MMGBSA) is one of the most commonly used end state approaches used for the calculation of the binding free energy towards computational drug design and screening studies. It is customary to break up the free energy into van der Waals, electrostatic, polar solvation (GB), and nonpolar solvation (SA) terms and then either correlate these terms with experiment or assign physical meaning to each term. Here, we demonstrate that this assumption of independent fitting coefficients for decomposed energy terms could be invalid. Through analytic derivation and large-scale molecular dynamics simulations, we show that (i) the protein and ligand Coulomb interaction energy and the GB solvation correction are almost perfectly collinear (R2[≥]0.99) reflecting their designed role as vacuum electrostatics plus solvent screening, and (ii) the van der Waals interaction and SA term likewise exhibit strong correlation, as both depend primarily on buried surface area. Interaction entropy and C2 entropy corrections are also found to be strongly dependent on underlying electrostatic fluctuations, further reinforcing redundancy. These findings hold both at the level of instantaneous trajectory fluctuations and when averaged across a diverse set of 139 protein-protein complexes and persist in both single-trajectory and three trajectory MMGBSA protocols. Our results caution against using decomposed MMGBSA terms as independent predictors in regression models and suggest instead combining correlated terms into effective polar, nonpolar, and entropic contributions. Our study provides a systematic diagnosis of collinearity in MMGBSA and highlights pathways toward more interpretable and statistically robust predictive modeling.

13
Characterizing the fragmentation of AlphaFold predictions

Sarti, E.; Cazals, F.

2026-06-05 bioinformatics 10.64898/2025.12.19.695436 medRxiv
Top 0.1%
1.8%
Show abstract

The Nobel prize winning program AlphaFold2 computes plausible structures of (well) folded proteins. The main quality assessment is based on the predicted Local Distance Difference Test (pLDDT), a per amino acid confidence score. To enhance quality assessment, we provide novel quantitative measures to identify coherent amino acid (a.a.) stretches along the sequence in terms of pLDDT values. These constructions, grounded in standard techniques from topological data analysis and combinatorics, provide a canonical framework for identifying regions along the protein backbone and analyzing their properties, such as their propensity for disorder and their consistency with a null model. The outcome of our analysis can readily be used to select reliable regions/domains within proteins whose pLDDT values span the entire pLDDT range.

14
Bayesian optimisation and graph-based rheology enable sequence-dependent modelling of DNA materials

Gadzekpo, A.; Hilbert, L.

2026-05-30 biophysics 10.64898/2026.05.27.728076 medRxiv
Top 0.1%
1.7%
Show abstract

Bridging molecular and emergent properties is essential for designing soft matter. Synthetic DNA materials are attractive in this context because their sequence design space supports a wide range of material properties. Targeted design of DNA materials is hindered by scale differences and manual exploration of vast design spaces. We address this challenge with a computational workflow that links sequence-level design to rheological material properties. Concretely, we use machine learning to parametrise scalable, DNA-sequence-aware simulations, which we then evaluate using graph-based rheology. In our example, we study materials composed of self-interacting, multivalent DNA nanostars assembled from single strands. Structure and flexibility of nanostars are quantified with nucleotide-level oxDNA simulations, enabling Bayesian optimisation of a more coarse-grained bead-spring model. The bead-spring model allows efficient simulation of network formation between nanostars, governed by hybridisation free energies, which are computed with oxDNA and NUPACK. Nanostar valency and network connectivity are translated into rheological material properties with a graph-based method that we extend to include hydrodynamic interactions, yielding good agreement with experimental reference data. We generalise our findings by analysing theoretical graph representations of DNA materials and show how machine learning can optimise sequence affinities to produce desired rheological responses. Our work illustrates how machine learning can bridge scales and automate coarse-graining to facilitate targeted design of DNA materials through sequence-property relationships. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/728076v1_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@20b8c6org.highwire.dtl.DTLVardef@42f843org.highwire.dtl.DTLVardef@b90119org.highwire.dtl.DTLVardef@1f72d66_HPS_FORMAT_FIGEXP M_FIG C_FIG

15
Bayesian-Steered Structure Prediction of Mechanical Biomolecules Using Twisted Diffusion

Klaus, C.; Sotomayor, M.

2026-05-13 bioinformatics 10.64898/2026.05.11.724187 medRxiv
Top 0.1%
1.7%
Show abstract

Deep learning approaches have revolutionized protein structure prediction. These tools are trained using experimental data and recapitulate reported conformations, but there is great interest in predicting conformations that may be functionally relevant although experimentally underrepresented. Since many modern structure prediction tools use generative artificial intelligence diffusion models, we reframe the search for alternative molecular conformations as that of sampling from a diffusion distribution conditioned using any arbitrary Bayesian likelihood. We implement a twisted diffusion sampler in Boltz-2 to sample this conditioned distribution and demonstrate the utility of this approach, which does not require any additional training of the neural network, by implementing a diffusion analog of steered molecular dynamics simulations applied to mechanical systems. We can reproduce predicted stretched states of fragments of DNA, the muscle protein titin, and the inner-ear protocadherin-15 protein, as well as open states of the MscL ion channel consistent with experimental results. We expect that steered structure predictions will help sample underrepresented and non-equilibrium conformations for many macromolecular systems.

16
Homology-aware cross-validation strategies for generalization assessment in RNA structure prediction

Bugnon, L.; Kulemeyer, G.; Gerard, M.; Di Persia, L.; Stegmayer, G.; Milone, D. H.

2026-06-29 bioinformatics 10.64898/2026.06.28.735057 medRxiv
Top 0.1%
1.7%
Show abstract

RNA secondary structure prediction is a fundamental challenge in bioinformatics, essential for understanding the functional roles of non-coding RNAs. Recently, deep learning models have transformed the field with impressive results, leading to critical discussions regarding the validity of current cross-validation strategies. On the one hand, traditional random partitioning yields overop-timistic results due to data leakage from uncontrolled homology. On the other hand, removing from the training set all sequences that exhibit even the slightest resemblance to the testing sequences penalizes learning-based methods by requiring generalization to completely out-of-distribution sequences. While it is very simple to remove sequences and retrain a machine learned model, it is very difficult to remove the experimental data used for parameter tuning and the sequences used for the development of classical thermodynamic methods. Thus, these methods often benefit from an implicit knowledge leakage. In this work we critically review existing cross-validation strategies for RNA secondary structure prediction: random splitting, clustering-based splitting, and leaving one RNA family out for testing. We analyze the advantages and limitations of each strategy, also expanding them towards the future directions to ensure fair comparisons across the full range of sequence similarities, with the same rigor for both classical and learning-based methods.

17
BioMetAll v2.0: Introducing Scores, Metal Discrimination, and Side-Chain Descriptors for Predicting Metal-Binding Sites in Proteins.

Marechal, J. D.; Fernandez Diaz, R.; Pena Losada, R.; Sanchez Aparicio, J. E.; Gao, W.; Alemany, M.

2026-07-12 bioinformatics 10.64898/2026.07.09.737562 medRxiv
Top 0.1%
1.7%
Show abstract

Predicting the location of metal-binding sites in proteins is crucial for fundamental biological questions and biotechnological applications. Over the past decade, the rise in metal-bound protein structures in the Protein Data Bank, combined with advanced statistical models such as deep learning, has accelerated the development of metal-binding site prediction tools. Several approaches are now available, offering high-quality benchmarks and predictive performance. Our initial development in this area is BioMetAll, whose first version was based on backbone pre-organization. Here, we introduce its second version, featuring two major updates: 1) metal-specific scoring functions and 2) prediction using backbone geometry alone or in combination with first coordination sphere descriptors. Apart from demonstrating metal sensitivity and yielding better benchmarking results, this new version allows the assessment of the influence of considering the metals first coordination sphere versus backbone pre-organization on how metallic species bind to proteins.

18
Capabilities, specificity gaps and training-data dependence of AlphaFold3 across diverse application areas

Follonier, O.; Liu, Y.; Campomanes, P.; Lafrenaye, L.; Racle, J.; Alvarez, D.; van Gerwen, J.; Heinzmann, R.; Jänes, J.; Kummelstedt, E.; Durairaj, J.; Gfeller, D.; Vanni, S.; Beltrao, P.

2026-07-13 bioinformatics 10.64898/2026.07.13.738147 medRxiv
Top 0.1%
1.7%
Show abstract

Structure prediction models have moved from single proteins to assemblies that include diverse biomolecules and their modifications. AlphaFold3 (AF3) and related models extended structural modelling via an all-atom framework, opening many new potential applications in structural biology. We evaluate how well the new capabilities of AF3 translate into application tasks in diverse areas: prediction of ubiquitinated protein structures, T-cell receptor (TCR)-epitope recognition, antibody-antigen complexes, protein-RNA and protein-lipid interactions. We find that, while AF3 can perform well in favourable settings, this performance is uneven across applications. In RNA-target predictions, the model confidence fails to separate genuine from decoy interaction partners and in several tasks accuracy depends on the presence of related complexes in the training set. Taken together, our assessment is more cautious than for AF2, whose gains in modelling monomers and complexes were clear and broadly generalisable. AF3s extension to new biomolecule types shows less consistent performance and generalisation. AF3 can be a powerful tool for hypothesis generation and prioritisation, but its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap. We expect that the benchmarks provided here will serve for testing of future developments in the structure prediction field.

19
MOFF2: A Transferable Coarse-Grained Protein Force Field for Predictive Condensate Simulations

Liu, S.; Zhang, Y.; Riveros, I.; Wang, C.; Zhang, B.

2026-06-10 biophysics 10.64898/2026.06.10.731384 medRxiv
Top 0.1%
1.7%
Show abstract

Coarse-grained protein force fields enable simulations of biomolecular systems at length and time scales that are difficult to access with atomistic models, but achieving transferability across folded, intrinsically disordered, and multidomain proteins remains challenging. A central difficulty is that one-bead-per-residue models must represent chemically specific residue interactions while also absorbing solvent-mediated and many-body effects into a simplified energy function. Here, we present MOFF2, a transferable coarse-grained protein force field that combines residue-pair-specific interactions with a density-dependent many-body potential. MOFF2 is optimized using a two-stage strategy: bottom-up parameter learning from heterogeneous reference ensembles followed by refinement against experimental conformational observables. The resulting model provides balanced performance across ordered proteins, intrinsically disordered proteins, and multidomain proteins, and predicts condensate saturation-concentration trends for A1-LCD variant systems. Analysis of the learned parameters reveals chemically interpretable interaction patterns and density-dependent effects that explain the models improved transferability. These results demonstrate that combining a generalized coarse-grained energy function with data-driven optimization can produce a practical and interpretable force field for protein conformational and condensate simulations.

20
fuzzyfold: a high-performance framework for stochastic RNA folding kinetics

Badelt, S.

2026-06-18 bioinformatics 10.64898/2026.06.17.732885 medRxiv
Top 0.1%
1.7%
Show abstract

The analysis of nucleic acid secondary structures is overwhelmingly dominated by methods that analyze the thermodynamic equilibrium distribution and which ignore all dynamic aspects of nucleic acid folding. Yet, there are numerous popular examples of nucleic acid folding that rely on kinetic models, such as RNA riboswitches or DNA strand displacement systems. Here, I am presenting fuzzyfold, a Rust-based software package for nucleic acid secondary structure analysis with an explicit focus on stochastic modeling. The framework introduces three-way and four-way shift moves with a biophysically motivated rate-model parameterization, and it is developed with an emphasis on both model flexibility and performance, e.g. allowing for the generation of single co-transcriptional trajectories for thousand-nucleotide long RNA molecules in just a few minutes. The main strength of the fuzzyfold package, however, is its focus on user and developer interfaces for long-term development. It provides easily installable command-line interfaces, e.g. for aggregating data from multiple parallel trajectories efficiently into an ensemble-level dynamic analysis. For developers, the code-base supports straight-forward substitution of thermodynamic and kinetic free-energy models, and a flexible library interface with Python bindings, enabling integration of individual components into custom computational workflows.